Papers with Language identification
STIL - Simultaneous Slot Filling, Translation, Intent Classification, and Language Identification: Initial Results using mBART on MultiATIS++ (2020.aacl-main)
Copied to clipboard
| Challenge: | Slot-filling, Translation, Intent classification, and Language identification (STIL) are tasks for multilingual Natural Language Understanding (NLU) . |
| Approach: | They propose to perform simultaneous slot filling and translation into a single output language (English in this case). |
| Outcome: | The proposed task performs better than the current state-of-the-art system for the languages tested, but with lower intent classification accuracy and lower slot F1 . |
An Open Dataset and Model for Language Identification (2023.acl-short)
Copied to clipboard
| Challenge: | Existing LID systems perform poorly on low-resource languages, causing 'representation washing', where the community is given a false view of the actual progress of low-source NLP. |
| Approach: | They propose a model which achieves a macro-average F1 score of 0.93 and a false positive rate of 0.033% across 201 languages, outperforming previous work. |
| Outcome: | The proposed model outperforms existing models and datasets on 201 languages and a false positive rate of 0.033%. |
AfroLID: A Neural Language Identification Tool for African Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | AfroLID is a neural LID toolkit for 517 African languages and varieties. |
| Approach: | They propose to exploit a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems to exploit AfroLID. |
| Outcome: | The proposed tool outperforms existing tools on the acutely under-served Twitter domain. |
Unsupervised Deep Language and Dialect Identification for Short Texts (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for identifying closely related short texts are unsupervised . however, performance is poor for unsupervised methods for short texts . |
| Approach: | They propose a method which can learn sentence embeddings and cluster assignments from short texts. |
| Outcome: | The proposed method outperforms state-of-the-art methods in supervised settings . it can learn sentence embeddings and cluster assignments from short texts . |
Subword-Level Language Identification for Intra-Word Code-Switching (N19-1)
Copied to clipboard
| Challenge: | Code-switching (CS) is a phenomenon of alternating between two or more languages in conversations . if at least one language is morphologically rich, a large number of words can be composed of morphemes from more than one language. |
| Approach: | They propose to extend the language identification task to the subword level by splitting mixed words while tagging each part with a language ID. |
| Outcome: | The proposed model outperforms the baseline on a Spanish–Wixarika and adapted German–Turkish datasets. |
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification (2026.eacl-long)
Copied to clipboard
| Challenge: | Low-resource languages and dialects remain difficult to identify and categorize accurately due to data in these languages and are limited to single-domain data. |
| Approach: | They propose a supervised contrastive learning approach to learn domain-invariant representations for low-resource languages by 3.2 percentage points while maintaining its performance for the high-resourced languages. |
| Outcome: | The proposed approach improves LID performance on out-of-domain data for low-resource languages by 3.2 percentage points while maintaining its performance for the high-resourced languages. |
Script-Agnosticism and its Impact on Language Identification for Dravidian Languages (2025.naacl-long)
Copied to clipboard
| Challenge: | a recent study shows that modern systems are script-dependent in language identification (langID) many languages are written in multiple writing systems, and script diversity is common in low-resource languages. |
| Approach: | They propose to learn script-agnostic representations using different strategies . they use word-level script randomization and script exposure to a language written in multiple scripts . |
| Outcome: | The proposed methods exploit script randomization and exposure to a language written in multiple scripts to improve language identification while maintaining competitive performance on naturally occurring text. |
Search Query Language Identification Using Weak Labeling (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has shown that language identification is a well-known task for natural language documents. |
| Approach: | They propose a search query language identification task that trains large-scale query-language pairs for training without loss of generalization. |
| Outcome: | The proposed model outperforms open domain model baselines by a large margin. |
FastSpell: The LangId Magic Spell (2024.lrec-main)
Copied to clipboard
| Challenge: | Language identification is a crucial component in the automated production of language resources. |
| Approach: | They propose a language identifier that combines fastText and Hunspell to give a second opinion before deciding which language to assign to a text. |
| Outcome: | The proposed language identifier is based on a pre-trained language identifier and a spell checker. |
GeezSwitch: Language Identification in Typologically Related Low-resourced East African Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | Low-resourced languages with similar typologies are often confused with each other in real-world applications such as machine translation, affecting the user’s experience. |
| Approach: | They propose to build a dataset for five typologically and phylogenetically related low-resourced East African languages using the Ge’ez script as a writing system. |
| Outcome: | The proposed dataset is built automatically from selected data sources, but also performed a manual evaluation to assess its quality. |
Language Variety Identification with True Labels (2024.lrec-main)
Copied to clipboard
Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice, Neha Kumari, Nishant Nair, Yash Mahesh Bangera
| Challenge: | Language identification datasets are compiled with the assumption that the gold label of each instance is determined by where texts are retrieved from. |
| Approach: | They present a human-annotated multilingual dataset for language variety identification . they use a model to train multiple models to discriminate between different languages . |
| Outcome: | The proposed dataset provides a reliable benchmark toward robust and fairer language variety identification systems. |
LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Currently, existing systems cannot accurately identify most of the world's 7000 languages due to lack of data and computational challenges. |
| Approach: | They propose a misprediction-resolution hierarchical model, LIMIT, that reduces error by 55% on a children's stories dataset and by 40% on 'fLORES-200' benchmark. |
| Outcome: | The proposed model reduces error by 55% on the MCS-350 and 40% on the FLORES-200 benchmarks. |
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data (2026.acl-long)
Copied to clipboard
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Akinyi Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob Van Der Goot, Lanwenn ar C’horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin L Rice, Azril Hafizi Amirudin, Jesujoba Oluwadara Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, null Akshata, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Rufaro Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger
| Challenge: | Language identification (LID) is a fundamental step in curating multilingual corpora. |
| Approach: | They introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. |
| Outcome: | The proposed benchmark covers 109 languages and shows that existing evaluations overestimate accuracy for many languages in the web domain. |